1 results listed
Visual Speech Recognition (VSR) is an important area of study. It focuses on turning silent visual speech cues into text. This research looks at a deep learning lip-reading model designed for the GRID dataset. The project works on getting video frames ready, creating visual animations and matching them with sound transcripts. A very interactive Streamlit interface helps users choose videos, see frames and create GIFs. Important contributions include improved methods for getting frames ready, new heatmap visuals for each frame and matching text with visuals for better understanding. This method shows a lot of promise in speech recognition fields. It shows that it can really help with accessibility and human-computer interaction.
International Conference on Advanced Technologies, Computer Engineering and Science
ICATCES
Adwyte Karandikar
Dr. Kaushalya Thopate
Ansh Sharma
Varad Adhyapak